Papers with Question answering

49 papers
Multi-Domain Multilingual Question Answering (2021.emnlp-tutorials)

Copied to clipboard

Challenge: Question answering (QA) is one of the most challenging tasks in natural language processing.
Approach: a tutorial examines the state-of-the-art approaches to multi-domain and multilingual QA . they introduce standard benchmarks and discuss out-of the-box training with open-domain QA systems .
Outcome: This tutorial aims to bridge the gap between open-domain and multilingual QA.
Reverse Question Answering: Can an LLM Write a Question so Hard (or Bad) that it Can’t Answer? (2025.naacl-short)

Copied to clipboard

Challenge: Question answering (QA) is a popular task, but we test both separately . a recent study found that LLMs are less accurate in numerical RQA than RQA .
Approach: We run 16 LLMs on QA and RQA with trivia questions/answers . they find question and answer types that lead to RQA errors and suggest improvements .
Outcome: The results show that LLMs are less accurate in RQA for numerical answers than RQA . RQA errors correlate with question difficulty and inversely correlate with answer frequencies .
A guide to the dataset explosion in QA, NLI, and commonsense reasoning (2020.coling-tutorials)

Copied to clipboard

Challenge: a tutorial aims to provide an up-to-date guide to the recent datasets . the target audience is the NLP practitioners who are lost in dozens of the recent data sets.
Approach: This tutorial provides an up-to-date guide to the recent datasets . it surveys old and new methodological issues with dataset construction .
Outcome: This tutorial aims to provide an up-to-date guide to the recent datasets . it surveys the old and new methodological issues with dataset construction .
TimelineQA: A Benchmark for Question Answering over Timelines (2023.findings-acl)

Copied to clipboard

Challenge: Existing question answering techniques for lifelogs do not provide accurate answers . augmented reality glasses have led to the creation of personal assistants .
Approach: They propose to use a benchmark to query lifelogs to find out what happened in real life . they find that extractive QA systems out-perform retrieval-augmented QA techniques .
Outcome: The proposed method outperforms state-of-the-art retrieval-augmented QA systems in atomic queries and multi-hop queries.
Question Answering in the Biomedical Domain (P19-2)

Copied to clipboard

Challenge: False positive questions require specific knowledge, common sense or a procedure due to ambiguity or the scope of the question.
Approach: False q is a question answering technique that uses natural language to find an answer . Falsity is based on a lexical gap and quality of answer spans .
Outcome: Using the proposed system, patients can self-diagnose without sacrificing quality of answer spans.
Fluent Response Generation for Conversational Question Answering (2020.acl-main)

Copied to clipboard

Challenge: Question answering (QA) is an important aspect of open-domain conversational agents, garnering specific research focus in the conversational QA subtask.
Approach: They propose a method for situating QA responses within a SEQ2SEQ NLG approach to generate fluent grammatical answer responses while maintaining correctness.
Outcome: The proposed model outperforms baseline CoQA and QuAC models in generating conversational responses.
Improving the Robustness of QA Models to Challenge Sets with Variational Question-Answer Pair Generation (2021.acl-srw)

Copied to clipboard

Challenge: Existing data augmentation methods for reading comprehension lack robustness to challenge sets whose distribution is different from that of training sets.
Approach: They propose a question-answer pair generation method that generates multiple diverse QA pairs from a paragraph to mitigate this problem.
Outcome: The proposed model improves the accuracy of 12 challenge sets and the in-distribution accuracy.
RobustQA: A Framework for Adversarial Text Generation Analysis on Question Answering Systems (2023.emnlp-demo)

Copied to clipboard

Challenge: Question answering (QA) systems have reached human-level accuracy, but they are not robust enough and vulnerable to adversarial examples.
Approach: They modified the attack algorithms widely used in text classification to fit them for QA systems.
Outcome: The proposed framework is the first open-source toolkit for investigating textual adversarial attacks in QA systems.
Retrieval, Re-ranking and Multi-task Learning for Knowledge-Base Question Answering (2021.eacl-main)

Copied to clipboard

Challenge: Existing work on question answering over knowledge bases limited the search space to a subset of KBs . a retrieval-and-rerank framework is used to access KB and rerank retrieved candidates with more powerful neural networks.
Approach: They propose to share a BERT encoder across all three sub-tasks and define task-specific layers on top of the shared layer.
Outcome: The proposed method improves accuracy and accuracy on the SimpleQuestions dataset and the FreebaseQA dataset.
Generalizing Question Answering System with Pre-trained Language Model Fine-tuning (D19-58)

Copied to clipboard

Challenge: Existing methods focus on improving in-domain performance, leaving open the question of how they can generalize to out-of-domain and unseen RC tasks.
Approach: They propose a multi-task learning framework that learns the shared representation across different tasks and builds on a large pre-trained language model and fine-tuned on multiple RC datasets.
Outcome: The proposed framework improves the BERT-Large baseline by 8.39 and 7.22 respectively.
ViMedAQA: A Vietnamese Medical Abstractive Question-Answering Dataset and Findings of Large Language Model (2024.acl-srw)

Copied to clipboard

Challenge: Existing abstractive question-answering datasets in Vietnamese are lacking .
Approach: They propose to introduce a Vietnamese abstractive question-answering corpus to address this gap . they propose to use Vietnamese abstractives to generate answers to questions .
Outcome: The proposed dataset examines the capability of large language models in the Vietnamese medical domain, including reasoning, memorizing and awareness of essential information.
MULTITAT: Benchmarking Multilingual Table-and-Text Question Answering (2025.findings-emnlp)

Copied to clipboard

Challenge: Existing TATQA datasets are limited to English, leading to drawbacks . existing datasets overlook challenges of multilingual TAT-QA and do not reflect real-world multilingual scenarios .
Approach: They propose a multilingual TATQA dataset that can be translated into 10 languages . they use data from 3 mainstream TATQ datasets and analyze the results .
Outcome: The proposed dataset outperforms other baselines by an average of 3.3 .
Fantastic Questions and Where to Find Them: FairytaleQA – An Authentic Dataset for Narrative Comprehension (2022.acl-long)

Copied to clipboard

Challenge: Existing QA datasets rarely distinguish fine-grained reading skills, such as the understanding of varying narrative elements.
Approach: They propose to use FairytaleQA to generate 10,580 questions based on 278 children-friendly stories to assess model's fine-grained learning skills.
Outcome: The proposed dataset consists of 10,580 questions derived from 278 children-friendly stories, covering seven types of narrative elements or relations.
Russian Jeopardy! Data Set for Question-Answering Systems (2022.lrec-1)

Copied to clipboard

Challenge: Question answering is one of the most common tasks in natural language processing . open-domain questions cover a wide range of topics and do not necessarily come in form of an actual question.
Approach: They describe a Russian question-like question set collected from the Russian analogue of Jeopardy! They observe its linguistic features and the related QA-task.
Outcome: The proposed data set includes 379,284 quiz-like questions with 29,375 from the Russian analogue of Jeopardy!
Towards more equitable question answering systems: How much more data do you need? (2021.acl-short)

Copied to clipboard

Challenge: Question answering datasets in English are relatively new, but lack of linguistic diversity in the field is a challenge.
Approach: They propose to use translation and cross-lingual transfer to produce QA systems in multiple languages to improve their performance.
Outcome: The proposed approaches take advantage of existing resources to produce QA systems in multiple languages.
ViDeBERTa: A powerful pre-trained language model for Vietnamese (2023.findings-eacl)

Copied to clipboard

Challenge: Existing models for Vietnamese that perform well on downstream tasks, such as Question answering, are based on Transformer.
Approach: They propose a pre-trained monolingual Vietnamese model with three versions . they fine-tune and evaluate the model on three important natural language downstream tasks, Part-of-speech tagging, Named-entity recognition, and Question answering.
Outcome: The proposed model outperforms the existing model on three important natural language downstream tasks, Part-of-speech tagging, Named-entity recognition, and Question answering.
AHP-Powered LLM Reasoning for Multi-Criteria Evaluation of Open-Ended Responses (2024.findings-emnlp)

Copied to clipboard

Challenge: Question answering (QA) tasks have been extensively studied in the field of natural language processing.
Approach: They propose a method that leverages large language models and the analytic hierarchy process to assess open-ended questions.
Outcome: The proposed method more closely aligns with human judgment compared to baselines on four datasets.
Learning to Collaborate for Question Answering and Asking (N18-1)

Copied to clipboard

Challenge: Question answering (QA) and question generation (QG) are closely related tasks.
Approach: They propose a training algorithm that generalizes both Generative Adversarial Network and Generating Domain-Adaptive Nets under the question answering scenario.
Outcome: The proposed training algorithm generalizes both Generative Adversarial Network (GAN) and Generating Domain-Adaptive Nets (GDAN) under the question answering scenario.
Logical Form Generation via Multi-task Learning for Complex Question Answering over Knowledge Bases (2022.coling-1)

Copied to clipboard

Challenge: Existing generation-based KBQA methods that translate natural language questions to executable logical forms are proving promising but noise introduced can lead to incorrect results.
Approach: They propose a Generation-based KBQA method that uses auxiliary information to enhance logical form generation by combining unseen KB items with novel combinations.
Outcome: The proposed method achieves state-of-the-art results on ComplexWebQuestions and WebQuestIONSSP datasets.
QA Domain Adaptation using Hidden Space Augmentation and Self-Supervised Contrastive Adaptation (2022.emnlp-main)

Copied to clipboard

Challenge: Question answering models often suffer from performance deterioration upon deployment .
Approach: They propose a self-supervised framework called QADA for QA domain adaptation . they propose to augment training QA samples with hidden space augmentation .
Outcome: The proposed framework improves on multiple target datasets over state-of-the-art methods.
ArcaneQA: Dynamic Program Induction and Contextualized Encoding for Knowledge Base Question Answering (2022.coling-1)

Copied to clipboard

Challenge: Existing ranking-based KBQA models struggle with flexibility in predicting complicated queries and have impractical running time.
Approach: They propose a new generation-based question answering on knowledge bases model that addresses both large search space and ambiguities in schema linking.
Outcome: The proposed model overcomes two intertwined challenges on popular KBQA datasets and is highly competitive and efficient.
Domain Adaptation for Question Answering via Question Classification (2022.coling-1)

Copied to clipboard

Challenge: Question answering systems often experience performance deterioration upon user-generated questions.
Approach: They propose a question classification framework to help QA domains adapt to different domains.
Outcome: The proposed framework improves on state-of-the-art datasets against multiple datasets.
UNIFIEDQA: Crossing Format Boundaries with a Single QA System (2020.findings-emnlp)

Copied to clipboard

Challenge: Question answering (QA) tasks have been posed using a variety of formats . a new study aims to develop specialized QA models that can be used to train QA systems .
Approach: They build a pre-trained question answering model that performs well across 19 QA datasets . they argue that format-specialized models can limit the ability to teach reasoning .
Outcome: a new model that trains on QA datasets performs on par with 8 models trained on individual datasets . a single model that trained on UNIFIEDQA performs well on 19 QA data .
Graph-Based Knowledge Integration for Question Answering over Dialogue (2020.coling-main)

Copied to clipboard

Challenge: Existing approaches for question answering over dialogue did not consider dialogue structure and background knowledge (e.g., relationships between speakers).
Approach: They propose a method which organizes a dialogue as a "relational graph" and uses edges to represent relationships between entities to encode multi-relations knowledge for reasoning.
Outcome: The proposed method is better at tackling complex questions requiring relational reasoning and defending adversarial attacks with distracting sentences.
Answering while Summarizing: Multi-task Learning for Multi-hop QA with Evidence Extraction (P19-1)

Copied to clipboard

Challenge: Question answering (QA) using textual sources for purposes such as reading comprehension has attracted much attention.
Approach: They propose a Query Focused Extractor model for evidence extraction and multi-task learning with the QA model.
Outcome: The proposed model achieves state-of-the-art evidence extraction score on hotpotQA and FEVER, which is a recognizing textual entailment task on a large textual database.
RAG-QA Arena: Evaluating Domain Robustness for Long-form Retrieval Augmented Question Answering (2024.emnlp-main)

Copied to clipboard

Challenge: Existing datasets for question answering based on retrieval augmented generation (RAG-QA) are either constructed using a single source corpus or consist of short extractive answers, which fall short of evaluating large language model (LLM) based RAG-QA systems on cross-domain generalization.
Approach: They propose a dataset that integrates short extractive answers from multiple documents into a single coherent narrative.
Outcome: The proposed dataset integrates short extractive answers from multiple documents into a single coherent narrative, covering 26K queries and large corpora across seven different domains.
SD-QA: Spoken Dialectal Question Answering for the Real World (2021.findings-emnlp)

Copied to clipboard

Challenge: Existing QA benchmarks do not account for errors that speech recognition models might introduce . evaluating production-ready QA systems on data that is not representative of real-world inputs is problematic .
Approach: They construct a multi-dialect, spoken QA benchmark on five languages with 68k audio prompts in 24 dialects from 255 speakers.
Outcome: The proposed model is based on 68k audio prompts in 24 dialects from 255 speakers.
Complex Question Answering on knowledge graphs using machine translation and multi-task learning (2021.eacl-main)

Copied to clipboard

Challenge: Existing approaches to question answering on knowledge graphs are based on a modularized sequential approach where errors in one module lead to the accumulation of errors in downstream modules.
Approach: They propose a multi-task BERT based Neural Machine Translation model to address these challenges.
Outcome: The proposed model can answer questions over a knowledge graph on one publicly available and one proprietary dataset.
A Neural Model for Joint Document and Snippet Ranking in Question Answering for Large Document Collections (2021.acl-long)

Copied to clipboard

Challenge: Question answering systems typically use pipelines that retrieve documents at finer text granularities.
Approach: They propose an architecture for document and snippet ranking that leverages intuition . they modified a natural questions dataset to test their model .
Outcome: The proposed model outperforms pipelines in document retrieval on biomedical data . the proposed model is competitive with the existing model, despite fewer parameters .
Few-shot In-context Learning on Knowledge Base Question Answering (2023.acl-long)

Copied to clipboard

Challenge: KB-BINDER enables few-shot in-context learning over knowledge base questions . KBQA is a difficult problem due to the heterogeneity of knowledge bases .
Approach: They propose a framework that enables few-shot in-context learning over KBQA tasks.
Outcome: The proposed framework can outperform state-of-the-art models on GraphQA and MetaQA datasets.
Multi-Row, Multi-Span Distant Supervision For Table+Text Question Answering (2023.acl-long)

Copied to clipboard

Challenge: Existing question answering systems for tables and linked text are relatively unexplored.
Approach: They propose a transformer-based question answering system that copes with distant supervision along both axes of the question and answer.
Outcome: The proposed system beats baselines for HybridQA and OTT-QA with best EM and F1 scores on a held out test set.
UniRPG: Unified Discrete Reasoning over Table and Text as Program Generation (2022.emnlp-main)

Copied to clipboard

Challenge: Existing methods for question answering using knowledge resources are mixed-of-experts and semantic parsing-based.
Approach: They propose a semantic-parsing-based approach to perform Unified discrete Reasoning over heterogeneous knowledge resources as Program Generation.
Outcome: The proposed approach improves interpretability and scalability over table and text . it achieves promising performance on the TAT-QA dataset without annotation .
PEDANTS: Cheap but Effective and Interpretable Answer Equivalence (2024.findings-emnlp)

Copied to clipboard

Challenge: Current short-form QA evaluations lack diverse styles of evaluation data and rely on expensive and slow LLMs.
Approach: They propose a rubric for machine QA that is more stable than an exact match and neural methods.
Outcome: The proposed evaluations improve on the existing short-form QA evaluations using the Trivia community.
Improving Time Sensitivity for Question Answering over Temporal Knowledge Graphs (2022.acl-long)

Copied to clipboard

Challenge: Temporal knowledge graphs record entity relations and when they occur in time . previous work fails to address time-related challenges such as time-order issues . paper proposes time-sensitive question answering framework to address these problems .
Approach: They propose a time-sensitive question answering framework that uses temporal KGs to answer natural language questions.
Outcome: The proposed framework outperforms the state-of-the-art on a new benchmark for question answering over temporal knowledge graphs.
Exploring Hybrid Question Answering via Program-based Prompting (2024.acl-long)

Copied to clipboard

Challenge: Existing approaches to question answering over heterogeneous data are limited due to large scale of information and organic coupling of heterogenous data.
Approach: They propose a program-based prompting framework for hybrid question answering tasks . it integrates various functions to perform hybrid information-seeking over data .
Outcome: The proposed framework surpasses baseline systems and achieves the best performance under the fewshot settings.
Can You Unpack That? Learning to Rewrite Questions-in-Context (D19-1)

Copied to clipboard

Challenge: Existing QA datasets lack key NLP problems like coreference and ellipsis resolution.
Approach: They propose a task of question-in-context rewriting to rewrite a context-dependent question into a self-contained question with the same answer.
Outcome: The proposed task is based on a dataset of 40,527 questions based in QuAC . it requires models to link questions together to resolve conversational dependencies .
ClarQ: A large-scale and diverse dataset for Clarification Question Generation (2020.acl-main)

Copied to clipboard

Challenge: Existing datasets hinder development of large-scale models capable of generating and utilising clarification questions.
Approach: They propose a bootstrapping framework that utilises a neural network architecture to classify clarification questions based on post-comment tuples extracted from stackexchange.
Outcome: The proposed framework aims to increase the accuracy of the classifier and increase recall of clarification questions by applying it to question-answering tasks.
MLQA: Evaluating Cross-lingual Extractive Question Answering (2020.acl-main)

Copied to clipboard

Challenge: Question answering (QA) models have shown rapid progress enabled by the availability of large, high-quality benchmark datasets.
Approach: They present a multi-way aligned extractive QA evaluation benchmark in 7 languages . they evaluate state-of-the-art cross-lingual models and machine-translation-based baselines .
Outcome: The proposed model is based on MLQA, which has over 12K instances in english and 5K in each other language.
GRAF: Graph Retrieval Augmented by Facts for Romanian Legal Multi-Choice Question Answering (2025.findings-acl)

Copied to clipboard

Challenge: Question answering systems have been used for various domains and languages.
Approach: They propose a novel approach for question answering (QA) that combines a dataset of Romanian legal questions with a CROL corpus of laws.
Outcome: The proposed approach achieves competitive results with generally accepted state-of-the-art methods and even exceeds them in most settings.
DRS: Deep Question Reformulation With Structured Output (2025.findings-acl)

Copied to clipboard

Challenge: Existing models like GPT-3 and Instruct-GPT lack the ability to reformulate unanswerable questions.
Approach: They propose a zero-shot method that combines the strengths of LLMs with a DFS-based algorithm to iteratively explore potential entity combinations and constrain outputs using predefined entities.
Outcome: The proposed method outperforms all baselines, including the GPT-3.5 model, on the unanswerable question reformulation task.
MusTQ: A Temporal Knowledge Graph Question Answering Dataset for Multi-Step Temporal Reasoning (2024.findings-acl)

Copied to clipboard

Challenge: Existing studies focus on fact-centered reasoning with limited attention to temporal reasoning.
Approach: They propose a new TKGQA dataset, MusTQ, which contains 666K multi-step temporal reasoning questions and a TKG.
Outcome: The proposed model achieves state-of-the-art multi-step temporal reasoning ability with entity-time attention mechanism and optimized temporal knowledge graph representation.
QAEval: Mixture of Evaluators for Question-Answering Task Evaluation (2025.acl-long)

Copied to clipboard

Challenge: Existing QA evaluation methods struggle with open-ended and unstructured responses.
Approach: They propose a hybrid framework that combines rule-based reliability with LLM-based adaptability to overcome these challenges.
Outcome: The proposed framework outperforms existing models like GPT-4o and Claude-3 in accuracy and cost.
LAD-RAG: Layout-aware Dynamic RAG for Visually-Rich Document Understanding (2026.acl-long)

Copied to clipboard

Challenge: Conventional retrieval-augmented generation (RAG) methods encode content in isolated chunks during ingestion, losing structural and cross-page dependencies, and retrieve a fixed number of pages at inference.
Approach: They propose a Layout-Aware Dynamic RAG framework that encodes content in isolated chunks during ingestion and retrieves a fixed number of pages at inference.
Outcome: Experiments on MMLongBench-Doc, LongDocURL, DUDE, and MP-DoxVQA show that LAD-RAG improves retrieval, achieving over 90% perfect recall on average without any top-k tuning, and outperforming baseline retrievers by up to 20% in recall at comparable noise levels.
Evaluation Paradigms in Question Answering (2021.emnlp-main)

Copied to clipboard

Challenge: Despite substantial overlap, subtle but significant distinctions exert an outsize influence on research . one paradigm values creating more intelligent QA systems, the other paradigm values building QA system that appeals to users.
Approach: They propose to use the Cranfield and Manchester paradigms to describe research working towards building human-like, intelligent QA systems.
Outcome: The proposed paradigms are based on the findings of two recent studies on question answering (QA) the Cranfield paradigm is not new, but the Manchester paradigm is christened as the most eclectic in QA .
Question Answering as Programming for Solving Time-Sensitive Questions (2023.emnlp-main)

Copied to clipboard

Challenge: Recent studies show that Large Language Models (LLMs) have shown remarkable intelligence in question answering.
Approach: They propose to reframe the Question Answering task as Programming to overcome this limitation by leveraging LLMs' superior ability in understanding both natural language and programming language.
Outcome: The proposed approach improves on time-sensitive question answering datasets by 14.5% over baselines.
SEER : A Knapsack approach to Exemplar Selection for In-Context HybridQA (2023.emnlp-main)

Copied to clipboard

Challenge: In-Context Learning with Large Language Models (LLMs) has shown great performance on reasoning tasks.
Approach: They propose a method for selecting a set of exemplars that is representative and diverse.
Outcome: The proposed method outperforms existing methods on FinQA and TAT-QA on hybrid questions.
Know the Known and the Unknown: Reasonable Answer Generation with Knowledge-Informed Citations (2026.acl-long)

Copied to clipboard

Challenge: Existing approaches focus on generating multi-level citations linked to specific references, making it verifiable and trustworthy.
Approach: They propose a new data construction pipeline and a benchmark to improve citation granularity and awareness of unknown information.
Outcome: The proposed model improves on the existing benchmark and data construction pipeline and provides citation granularity and awareness of unknown information.
Recursive Question Understanding for Complex Question Answering over Heterogeneous Personal Data (2025.findings-acl)

Copied to clipboard

Challenge: a novel method for question answering over mixed sources, like text and tables, has been developed for question-answering . personal information is a prominent case of such heterogeneous data, such as calendar entries, workout statistics, shopping records, streaming history, and more.
Approach: They propose a method that creates an executable operator tree for a given question . they use recursive decomposition to decompose a question into an operator tree .
Outcome: The proposed method outperforms methods based on verbalization or translation . it can be executed on user devices and yields a traceable answer .
VoiceBBQ: Investigating Effect of Content and Acoustics in Social Bias of Spoken Language Model (2025.emnlp-main)

Copied to clipboard

Challenge: Due to the nature of speech modality, social bias in Spoken Language Models (SLMs) can emerge from two distinct sources: 1) content aspect and 2) acoustic aspect.
Approach: They propose a dataset that measures social bias by presenting ambiguous or disambiguated contexts followed by questions that may elicit stereotypical responses.
Outcome: The proposed dataset converts every BBQ context into controlled voice conditions, enabling per-axis accuracy, bias, and consistency scores comparable to the original text benchmark.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations